Goto

Collaborating Authors

 safety and trustworthiness


JT-Safe: Intrinsically Enhancing the Safety and Trustworthiness of LLMs

arXiv.org Artificial Intelligence

The hallucination and credibility concerns of large language models (LLMs) are global challenges that the industry is collectively addressing. Recently, a significant amount of advances have been made on post-training and inference techniques to mitigate these challenges. However, it is widely agreed that unsafe and hallucinations of LLMs intrinsically originate from pre-training, involving pre-training data and the next-token prediction learning mechanism. In this paper, we focus on enhancing pre-training data to improve the trustworthiness and safety of LLMs. Since the data is vast, it's almost impossible to entirely purge the data of factual errors, logical inconsistencies, or distributional biases. Moreover, the pre-training data lack grounding in real-world knowledge. Each piece of data is treated as a sequence of tokens rather than as a representation of a part of the world. To overcome these issues, we propose approaches to enhancing our pre-training data with its context in the world and increasing a substantial amount of data reflecting industrial scenarios. We argue that most source data are created by the authors for specific purposes in a certain spatial-temporal context. They have played a role in the real world. By incorporating related world context information, we aim to better anchor pre-training data within real-world scenarios, thereby reducing uncertainty in model training and enhancing the model's safety and trustworthiness. We refer to our Data with World Context as DWC. We continue pre-training an earlier checkpoint of JT-35B-Base with 1.5 trillion of DWC tokens. We introduce our post-training procedures to activate the potentials of DWC. Compared with the Qwen model of a similar scale, JT-Safe-35B achieves an average performance improvement of 1.79% on the Safety and Trustworthy evaluation benchmarks, while being pretrained with only 6.2 trillion tokens.


From Reddit to Generative AI: Evaluating Large Language Models for Anxiety Support Fine-tuned on Social Media Data

arXiv.org Artificial Intelligence

The critical shortage of mental health services due to workforce limitations and logistical barriers, especially in underserved areas designated by the Health Resources & Services Administration (HRSA) 1, highlights the urgent need for accessible and scalable solutions. Traditional services often fail to address the diverse needs of individuals experiencing anxiety, prompting many, especially younger populations, to seek alternative emotional and psychological support online. While digital platforms offer immediate access, unregulated online interactions, including those with generative AI, may disseminate misleading information or inappropriate advice, potentially exacerbating anxiety symptoms (Tobias & Ito, 2021). Despite the great potential of generative AI to supplement mental health services, its deployment poses potentially significant risks. Unlike clinical practitioners, LLMs are not inherently equipped to manage emotionally complex or vulnerable conversations, which are critical to therapeutic relationships that create positive clinical outcomes (Rogers, 1957; Wampold, 2015).


Investigating the Impact of Quantization Methods on the Safety and Reliability of Large Language Models

arXiv.org Artificial Intelligence

Large Language Models (LLMs) have emerged as powerful tools for addressing modern challenges and enabling practical applications. However, their computational expense remains a significant barrier to widespread adoption. Quantization has emerged as a promising technique to democratize access and enable low resource device deployment. Despite these advancements, the safety and trustworthiness of quantized models remain underexplored, as prior studies often overlook contemporary architectures and rely on overly simplistic benchmarks and evaluations. To address this gap, we introduce OpenSafetyMini, a novel open-ended safety dataset designed to better distinguish between models. We evaluate 4 state-of-the-art quantization techniques across LLaMA and Mistral models using 4 benchmarks, including human evaluations. Our findings reveal that the optimal quantization method varies for 4-bit precision, while vector quantization techniques deliver the best safety and trustworthiness performance at 2-bit precision, providing foundation for future research.


Towards AI-$45^{\circ}$ Law: A Roadmap to Trustworthy AGI

arXiv.org Artificial Intelligence

Ensuring Artificial General Intelligence (AGI) reliably avoids harmful behaviors is a critical challenge, especially for systems with high autonomy or in safety-critical domains. Despite various safety assurance proposals and extreme risk warnings, comprehensive guidelines balancing AI safety and capability remain lacking. In this position paper, we propose the \textit{AI-\textbf{$45^{\circ}$} Law} as a guiding principle for a balanced roadmap toward trustworthy AGI, and introduce the \textit{Causal Ladder of Trustworthy AGI} as a practical framework. This framework provides a systematic taxonomy and hierarchical structure for current AI capability and safety research, inspired by Judea Pearl's ``Ladder of Causation''. The Causal Ladder comprises three core layers: the Approximate Alignment Layer, the Intervenable Layer, and the Reflectable Layer. These layers address the key challenges of safety and trustworthiness in AGI and contemporary AI systems. Building upon this framework, we define five levels of trustworthy AGI: perception, reasoning, decision-making, autonomy, and collaboration trustworthiness. These levels represent distinct yet progressive aspects of trustworthy AGI. Finally, we present a series of potential governance measures to support the development of trustworthy AGI.


ST-WebAgentBench: A Benchmark for Evaluating Safety and Trustworthiness in Web Agents

arXiv.org Artificial Intelligence

Recent advancements in Web agents have introduced novel architectures and benchmarks showcasing progress in autonomous web navigation and interaction. However, most existing benchmarks prioritize effectiveness and accuracy, overlooking factors like safety and trustworthiness which are essential for deploying web agents in enterprise settings. We present STWebAgentBench, a benchmark designed to evaluate web agents safety and trustworthiness across six critical dimensions, essential for reliability in enterprise applications. This benchmark is grounded in a detailed framework that defines safe and trustworthy (ST) agent behavior. Our work extends WebArena with safety templates and evaluation functions to assess safety policy compliance rigorously. We introduce the Completion Under Policy to measure task success while adhering to policies, alongside the Risk Ratio, which quantifies policy violations across dimensions, providing actionable insights to address safety gaps. Our evaluation reveals that current SOTA agents struggle with policy adherence and cannot yet be relied upon for critical business applications. We open-source this benchmark and invite the community to contribute, with the goal of fostering a new generation of safer, more trustworthy AI agents. All code, data, environment reproduction resources, and video demonstrations are available at https://sites.google.com/view/st-webagentbench/home.


Standardization Trends on Safety and Trustworthiness Technology for Advanced AI

arXiv.org Artificial Intelligence

Artificial intelligence (AI) technology has been evolving more rapidly over the past decade. With new ML models, data sources, and increased computational power, AI researchers have developed AI technologies that can understand language, recognize and create images and videos, program, and make scientific inferences. Recent advances in advanced AI technologies have evolved beyond traditional narrow domain AI to approximate or exceed artificial general intelligence (AGI) based on large language models (LLMs) or foundation models (FMs). These advanced AI systems are performing at or above human levels in complex problem solving, sophisticated natural language processing, and multi-domain tasks, and have the potential to revolutionize a wide range of fields, including science, industry, healthcare, and education. They are already surpassing human capabilities in certain task domains, such as Go, strategy games, and protein folding prediction [1] [2]. For these reasons, concerns about the safety and trustworthiness of advanced AI are growing rapidly alongside its development. The increasing complexity and autonomy of advanced AI systems is raising concerns that they could lead to new forms of safety and security risks, such as (1) uncontrollability, (2) conflicts with human values in ethical decision-making, (3) long-term socioeconomic impacts, and (4) safety assurance. In response, international standardization efforts are underway to ensure the safety and trustworthiness of advanced AI. By developing internationally agreed technical standards, efforts are being made to apply consistent safety and trustworthiness criteria to the development and use of advanced AI systems and minimize potential risks.


The often underestimated piece to successful Artificial Intelligence

#artificialintelligence

The first generation of AI has picked up on human biases. Among many disturbing cases of biased AI systems resulting in discriminatory outcomes, the most heart-breaking ones were cases involving unfair elongation of prison sentence, unfair credit card decision, and home appraisal outcomes. So, how does bias get into AI systems? While this is by no means an excuse, it does point to the key problem -- almost no focus was given to ensuring the moral, social, and responsible aspect of AI- often termed Ethical AI. A 2019 Gartner study reported that by 2022, 30% of the companies will invest in explainable ethical AI, from almost none in 2019.


Safety and Trustworthiness of Deep Neural Networks: A Survey

arXiv.org Artificial Intelligence

In the past few years, significant progress has been made on deep neural networks (DNNs) in achieving human-level intelligence on several long-standing tasks. With broader deployment of DNNs on various applications, the concerns on its safety and trustworthiness have been raised, particularly after the fatal incidents of self-driving cars. Research to address these concerns is very active, with many papers released in the past few years. This survey paper is to conduct a review of the current research efforts on making DNNs safe and trustworthy, by focusing on four aspects, i.e., verification, testing, adversarial attack and defence, and interpretability. In total, we surveyed 178 papers, most of which were published in the most recent two years, i.e., 2017 and 2018.